Papers with preprocessing step
Neural Mention Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | Mention detection is an important preprocessing step for downstream applications such as NER and coreference resolution. |
| Approach: | They propose and compare three approaches to mention detection using ELMO embeddings and a biaffine classifier. |
| Outcome: | The proposed model outperforms state-of-the-art models on the GENIA corpora and improves on mention recall. |
CoQAR: Question Rewriting on CoQA (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing systems that ask questions in a conversational context may have contextual dependencies that make the understanding difficult. |
| Approach: | They propose to rewrite questions into an out-of-context form to facilitate understanding . they propose to use this form to train and evaluate conversational question answering models . |
| Outcome: | The proposed model can be used in the supervised learning of three tasks: question paraphrasing, question rewriting and conversational question answering. |
Multilingual Email Zoning (2021.eacl-srw)
Copied to clipboard
| Challenge: | Existing literature on email zoning is mainly limited to English . however, it is possible to discern a level of formal organization in the way most emails are formed. |
| Approach: | They propose a multilingual email zoning benchmark based on a language agnostic sentence encoder and a new model that uses a biLSTM with a CRF to classify each sentence into an email zone. |
| Outcome: | The proposed model is competitive with current English benchmarks and reached state-of-the-art performance in English. |
Discourse-Based Sentence Splitting (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Sentence splitting is a key component of sentence simplification and has been shown to help human comprehension. |
| Approach: | They propose to use a discourse connective to generate a sentence that is shorter than the input text. |
| Outcome: | The proposed models outperform end-to-end models in learning the various ways of expressing a discourse relation but generate text that is less grammatical. |
Goodwill Hunting: Analyzing and Repurposing Off-the-Shelf Named Entity Linking Systems (2021.naacl-industry)
Copied to clipboard
| Challenge: | Named entity linking (NEL) is a preprocessing step in commercial systems . a small organization or individual could use an off-the-shelf system to accomplish the same objectives . |
| Approach: | They examine how to repurpose off-the-shelf NEL systems to correct sport-related errors. |
| Outcome: | The proposed model can improve sports question-answering accuracy by 25% . the proposed model is based on the best available model . |
A Simple Approach for Handling Out-of-Vocabulary Identifiers in Deep Learning for Source Code (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to handle out-of-vocabulary identifiers are not suitable for source code processing. |
| Approach: | They propose a method to handle out-of-vocabulary identifiers by identifies anonymization . they show that the method significantly improves the performance of the Transformer . |
| Outcome: | The proposed method significantly improves the performance of the Transformer in two code processing tasks. |
Increasing In-Class Similarity by Retrofitting Embeddings with Demographic Information (D18-1)
Copied to clipboard
| Challenge: | a new method for text classification ignores strong non-linguistic similarities like homophily . authors are typically represented via their linguistic profiles, i.e. information avail-able in the text . |
| Approach: | They use homophily cues to retrofit text-based author representations with non-linguistic information and introduce a trade-off parameter. |
| Outcome: | The proposed method improves on two author-attribute prediction tasks with large labels. |
Multilingual Normalization of Temporal Expressions with Masked Language Models (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for normalizing temporal expressions are rule-based, which severely limits the applicability in multilingual settings. |
| Approach: | They propose a neural method for normalizing temporal expressions based on masked language modeling and a slot-based prediction scheme for context-independent representations. |
| Outcome: | The proposed method outperforms existing rule-based methods in many languages and in particular, for low-resource languages with performance improvements of up to 33 F1 on average compared to the state of the art. |
Evaluating Historical Text Normalization Systems: How Well Do They Generalize? (N18-2)
Copied to clipboard
| Challenge: | Historical text normalization systems aim to convert historical wordforms to their modern equivalents . many of these systems have been developed and tested on a single language . |
| Approach: | They propose to use a nave baseline system to evaluate historical text normalization systems . they show that the models generalize well to unseen words in tests on five languages . |
| Outcome: | The proposed models generalize well to unseen words on five languages, but provide no clear benefit over the nave baseline. |
DR-BiLSTM: Dependent Reading Bidirectional LSTM for Natural Language Inference (N18-1)
Copied to clipboard
Reza Ghaeini, Sadid A. Hasan, Vivek Datla, Joey Liu, Kathy Lee, Ashequl Qadir, Yuan Ling, Aaditya Prakash, Xiaoli Fern, Oladimeji Farri
| Challenge: | Existing approaches to natural language inference rely on simple reading mechanisms for independent encoding of the premise and hypothesis. |
| Approach: | They propose a novel bidirectional dependent reading network to efficiently model the relationship between a premise and a hypothesis during encoding and inference. |
| Outcome: | The proposed model outperforms existing methods by a considerable margin on the Stanford Natural Language Inference (SNLI) dataset. |
Fast WordPiece Tokenization (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for tokenization of text are not efficient, but they are based on Aho-Corasick's algorithm. |
| Approach: | They propose an efficient algorithm for WordPiece tokenization using a longest-match-first strategy . they propose an algorithm whose tokenization complexity is strictly O(n) |
| Outcome: | The proposed method is 8.2x faster than HuggingFace Tokenizers and 5.1x faster on average for general text tokenization. |
Subword Segmental Machine Translation: Unifying Segmentation and Target Sentence Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Subword segmenters are used in neural machine translation, but are not used in high-resource settings. |
| Approach: | They propose a subword segmental machine translation (SSMT) that unifies subword and MT in a single trainable model. |
| Outcome: | The proposed model improves chrF scores for morphologically rich agglutinative languages and is more robust on a test set constructed for evaluating morphology generalisations. |
Parsivar: A Language Processing Toolkit for Persian (L18-1)
Copied to clipboard
| Challenge: | a preprocessing step is required to convert text into a standard format for NLP tasks. |
| Approach: | They propose a Persian preprocessing toolkit that performs various kinds of activities . they use a plagiarism detection system to exploit the proposed toolkit . |
| Outcome: | The proposed tool outperforms available Persian preprocessing tools by about 8 percent in terms of F1 . the proposed toolkit performs normalization, space correction, tokenization, stemming, parts of speech tagging and shallow parsing tasks. |
From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding (2023.acl-long)
Copied to clipboard
| Challenge: | Current models for natural language understanding require a preprocessing step to convert raw text into discrete tokens. |
| Approach: | They propose a hierarchical open-vocabulary language model that adopts a shallow Transformer architecture to learn word representations from their characters and a deep inter-word Transformer module that contextualizes each word representation by attending to the entire word sequence. |
| Outcome: | The proposed model outperforms baselines on various downstream tasks and is robust to textual corruption and domain shift. |
Harnessing Pre-Trained Neural Networks with Rules for Formality Style Transfer (D19-1)
Copied to clipboard
| Challenge: | Existing studies normalize informal sentences with rules, but they introduce noise if we use them in a naive way. |
| Approach: | They propose to harness rules into a state-of-the-art neural network that is typically pretrained on massive corpora. |
| Outcome: | The proposed method can be used to generate a state-of-the-art on a small dataset. |
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Subword tokenization approaches misalign with linguistic structure and waste capacity across languages and domains. |
| Approach: | They argue for a context-aware framework that integrates tokenizer and model co-design . they argue that tokenization should be treated as a core design problem, not an afterthought . |
| Outcome: | The proposed framework integrates tokenizer and model co-design, guided by linguistic, domain, and deployment considerations. |
Subword Segmental Language Modelling for Nguni Languages (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Subword segmentation is a standard practice in NLP, but is viewed as a preprocessing step for low-resource languages with complex morphologies. |
| Approach: | They propose a subword segmental language model that learns how to segment words while being trained for autoregressive language modelling. |
| Outcome: | The proposed model outperforms existing models on unsupervised morphological segmentation and outperfies standard subword segmenters on all 4 languages. |
Curating Datasets for Better Performance with Example Training Dynamics (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to improve data quality but rely on data quantity to improve performance are not effective. |
| Approach: | They propose a method for weighing the relative importance of examples in a dataset based on their Example Training dynamics (ETD) they propose an active learning approach for computing ETD during training rather than as a preprocessing step. |
| Outcome: | The proposed method can be used to improve performance in in-distribution and out-of-distortion testing. |
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Morphologically rich languages are notoriously challenging to process for downstream NLP applications. |
| Approach: | They propose a pretrained model for NLP applications involving the morphologically rich language Sanskrit that outperforms previous models by a considerable margin. |
| Outcome: | The proposed model outperforms tokenized models on established Sanskrit word segmentation tasks and matches the current best lexicon-based model. |